Papers by Mohammad Aflah Khan
Probing Critical Learning Dynamics of PLMs for Hate Speech Detection (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies on pretrained language models (PLMs) for hate speech detection have not investigated how their performance is affected by pretraining and finetuning. |
| Approach: | They propose to compare pretrained language models, evaluate their seed robustness, finetuning settings, and the impact of pretraining data collection time. |
| Outcome: | The proposed models show that they are more robust than other models and that they have a better chance of performing better than domain-specific models. |
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective (2026.acl-long)
Copied to clipboard
Bishwamittra Ghosh, Soumi Das, Till Speicher, Qinyuan Wu, Mohammad Aflah Khan, Deepak Garg, Krishna P. Gummadi, Evimaria Terzi
| Challenge: | Prior studies comparing FT and ICL have yielded mixed and inconclusive results due to inconsistent experimental setups. |
| Approach: | They propose a formal language learning task with precise language boundaries, controlled string sampling, and no data contamination to enable a rigorous comparison. |
| Outcome: | The proposed task offers precise language boundaries, controlled string sampling, and no data contamination. |
QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs (2025.coling-main)
Copied to clipboard
| Challenge: | QUENCH is a text-based English quizzing benchmarking system for large language models (LLMs). |
| Approach: | They propose a text-based English Quizzing Benchmark manually curated from YouTube quiz videos. |
| Outcome: | The proposed system assesses the world knowledge and deduction capabilities of large language models via a zero-shot, open-domain quizzing setup. |
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability (2025.emnlp-demos)
Copied to clipboard
Mohammad Aflah Khan, Ameya Godbole, Johnny Wei, Ryan Yixiang Wang, James Flemings, Krishna P. Gummadi, Willie Neiswanger, Robin Jia
| Challenge: | Existing workflows for pretraining large language models are cumbersome, fragmented and inaccessible. |
| Approach: | They propose an open-source library for editing, inspection, and analysis of large language model datasets. |
| Outcome: | TokenSmith is an open-source library for editing, inspection, and analysis of large language model datasets. |